Back

BMC Medical Genomics

Springer Science and Business Media LLC

All preprints, ranked by how well they match BMC Medical Genomics's content profile, based on 50 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Uncovering hidden gene-trait patterns through biclustering analysis of the UK Biobank

Pividori, M.; Sadeeq, S.; Krishnan, A.; Stranger, B. E.; Gignoux, C.

2024-11-11 bioinformatics 10.1101/2024.11.08.622657 medRxiv
Top 0.1%
39.2%
Show abstract

The growing availability of genome-wide association studies (GWAS) and large-scale biobanks provides an unprecedented opportunity to explore the genetic basis of complex traits and diseases. However, with this vast amount of data comes the challenge of interpreting numerous associations across thousands of traits, especially given the high polygenicity and pleiotropy underlying complex phenotypes. Traditional clustering methods, which identify global patterns in data, lack the resolution to capture overlapping associations relevant to subsets of traits or genes. Consequently, there is a critical need for innovative analytic approaches capable of revealing local, biologically meaningful patterns that could advance our understanding of trait comorbidities and gene-trait interactions. Here, we applied BiBit, a biclustering algorithm, to transcriptome-wide association study (TWAS) results from PhenomeXcan, a large resource of gene-trait associations derived from the UK Biobank. BiBit allows simultaneous grouping of traits and genes, identifying biclusters that represent local, overlapping associations. Our analyses uncovered biologically interpretable patterns, including asthma-related biclusters enriched for immune-related gene sets, connections between eye traits and blood pressure, and associations between dietary traits, high cholesterol, and specific loci on chromosome 19. These biclusters highlight gene-trait relationships and patterns of trait co-occurrence that may otherwise be obscured by traditional methods. Our findings demonstrate that biclustering can provide a nuanced view of the genetic architecture of complex traits, offering insights into pleiotropy and disease mechanisms. By enabling the exploration of complex, overlapping patterns within biobank-scale datasets, this approach provides a valuable framework for advancing research on genetic associations, comorbidities, and polygenic traits.

2
A Weights-based variant ranking pipeline for familial complex disorders

Ralli, S.; Vira, T.; Robles-Espinoza, C. D.; Adams, D. J.; Brooks-Wilson, A. R.

2023-08-16 genomics 10.1101/2023.08.14.553248 medRxiv
Top 0.1%
26.3%
Show abstract

Identifying genetic susceptibility factors for complex disorders remains a challenging task. We have developed a weights-based pipeline to prioritize variants and genes in collections of small and large pedigrees where genetic heterogeneity is likely, but biological commonalities are plausible. The Weights-based vAriant Ranking in Pedigrees (WARP) pipeline prioritizes variants using 5 weights: disease incidence rate, number of cases in a family, genome fraction shared amongst cases in a family, allele frequency and variant deleteriousness. Weights, except for the population allele frequency weight, are normalized between 0 to 1. Weights are combined multiplicatively to produce family-specific-variant weights that are then averaged across all families in which the variant is observed to generate a multifamily weight. Sorting multifamily weights in descending order creates a ranked list of variants and genes for further investigation. WARP was validated using familial melanoma sequence data from the European Genome-phenome Archive. The pipeline identified variation in known germline melanoma genes POT1, MITF and BAP1 in 4 out of 13 families (31%). Analysis of the other 9 families identified several interesting genes, some of which might have a role in melanoma. WARP provides an approach to identify disease predisposing genes in studies with small and large pedigrees.

3
The unreachable genomic profiling of complex diseases: genotype missingness matters

Abad-Grau, M. M.

2025-07-31 bioinformatics 10.1101/2025.07.27.667026 medRxiv
Top 0.1%
22.8%
Show abstract

The problem of building genome-wide predictors of individual risk to complex diseases seems to be more challenging than it was thought when the first human genome was sequenced on 2003. We have build different enhanced genetic risk predictors from genome-wide data and different complex diseases, making use of haplotypes accurately ascertained from family trios. We confirmed the widely known inability to accurately predict individual risk to complex diseases returned by the state-of-the-art genome-wide predictors. This result is mainly due to the small effect that most genetic variants have in a disease. We also found out that rates of missing genotypes were usually too high for these small-effect variants, as we could force missing imputation in such a tricky way that we would build highly accurate predictors, either by using our own design or the state-of-the-art genetic predictors. We observed that unknown genotypes were not missing at random but related to disease affectation, with more missing genotypes in affected than in non-affected individuals. We were not able to find a way to accurately reduce missing rates to correctly improve accuracy, but we identified a common pattern of missing data in multiple sclerosis, asthma and autism that makes us think that other complex diseases could behave the same way. Because (1) missing rates are high enough to completely change risk prediction due to the small-effect of most of the variants, and (2) there are more missing genotypes in affected than in non affected individuals, we conclude that perhaps the widely known defeat in genomic profiling for complex diseases may be solved by looking closer to the way current genotyping technologies handle genetic variants that may be rare in reference panels but have some effect in a given complex disease.

4
A phenotype-specific framework for identifying the eye abnormalities causative nonsynonymous-variants

Liu, H.; Dang, X.; Guan, L.; Tian, C.; Zhang, S.; Ye, C.; M. Tellier, L. C. A.; Chen, F.; Yang, H.; Sun, H.; Wu, J.; Zhang, J.

2020-04-13 bioinformatics 10.1101/2020.04.13.038059 medRxiv
Top 0.1%
22.7%
Show abstract

The most important role of variant pathogenicity predictors is to identify the disease-phenotype causative variant in studying monogenic diseases. In the last decade, machine-learning based predictors exhibited a relatively accurate performance for distinguishing the pathogenic variants and contributed a significant role for all disease-spectrums. Yet, few predictors can investigate the phenotypic significance of variants. Here we presented a phenotype-specific framework aimed to directly point out the phenotypic significance of predicted candidates, and showed its advancing performance in eye abnormalities. By training on eye-abnormalities causative variants, our method presented 96.2% accuracy, 96.1% precision, 93.4% recall for pathogenicity identification. Inconsistent with the modeling performance, identifying the single phenotype-causative variant from various sequencing variants is challenging for all predictors. Underlying the phenotype-oriented, our method significantly promoted the precision and reduced the cost for identifying the single causative variant from thousands of candidates. These advances highlight the significance of the phenotype-specific training method for studying disease.

5
A Pilot Meta-research on Evolving Evidence Behind Genetic Variant (Re)Classification

Ma, H.; Xu, Z.; Chung, W.; Weng, C.; Peng, Y.

2025-04-09 health informatics 10.1101/2025.04.07.25325116 medRxiv
Top 0.1%
22.7%
Show abstract

Variant classification and reclassification are fundamental to advancing precision medicine. This study focuses on the reclassifications of variants of uncertain significance (VUS) in BRCA1 and BRCA2 genes. By analyzing 162 unique cited publications supporting VUS reclassifications, we examined the accuracy, completeness, and currency of citations to these publications. Our findings reveal missing or inadequate evidence for reclassifications, as well as temporally misaligned citations and ClinVar submissions. Furthermore, we observed patterns in the cited studies, including the use of classification recommendations, genetic mechanisms, computational tools, and diverse population studies. This study underscores the need for stronger evidence supporting reclassifications and greater inclusion of diverse populations to optimize genomic variant reclassification and clinical decision-making.

6
Software as a Service for the Genomic Prediction of Complex Diseases.

Bolli, A.; Di Domenico, P.; Botta, G.

2019-09-11 genomics 10.1101/763722 medRxiv
Top 0.1%
22.0%
Show abstract

In the last decade the scientific community witnessed a large increase in Genome-Wide Association Study sample size, in the availability of large Biobanks and in the improvements of statistical methods to model genomes features. This have paved the way for the development of new prediction medicine tools that use genomic data to estimate disease risk. One of these tools is the Polygenic Risk Score (PRS), a metric that estimates the genetic risk of an individual to develop a disease, based on a combination of a large number of genetic variants.\n\nUsing the largest prospective genotyped cohort available to date, the UK Biobank, we built a new PRS for Coronary Artery Disease (CAD) and assessed its predictive performances along with two additional PRS for Breast Cancer (BC), and Prostate Cancer (PC). When compared with previously published PRS, the newly developed PRS for CAD displayed higher AUC and positive predictive value. PRSs were able to stratify disease risks from 1.34% to 25.7% (CAD in men), from 0.26% to 8.62% (CAD in women), from 1.6% to 24.6% (BC), and from 1.4% to 24.3% (PC) in the lowest and highest percentiles, respectively. Additionally, the three PRSs were able to identify the 5% of the UK Biobank population with a relative risk for the diseases at least 3 times higher than the average.\n\nFamily history is a well recognised risk factor of CAD, BC, and PC and it is currently used to identify individuals at high risk of developing the diseases. We show that individuals with family history can have completely different disease risks based on PRS stratification: from 2.1% to 33% (CAD in men), from 0.56% to 10% (CAD in women), from 2.3% to 35.8% (BC), and from 1.0% to 34.0% (PC) in the lowest and highest percentiles, respectively. Additionally, the PRSs demonstrated higher predictive performance (AUCs (including age) CAD: 0.81, PC: 0.80, and BC: 0.68) than family history (AUCs (including age) CAD: 0.79, PC: 0.73, and BC: 0.61) in predicting the onset of diseases.\n\nHyperlipidemia is well known to be associated with higher CAD risk, but a predictive performance comparison between each lipoprotein and CAD PRS has never been assessed. PRS shows higher discrimination capacity and Odds ratio per Standard deviation than LDL, HDL, total cholesterol-HDL ratio, ApoA, ApoB, ApoB-ApoA ratio, and Lipoprotein(a). Comparing the empirical risk distribution between PRS and each lipoprotein, we show that lipoprotein thresholds, currently used in clinical practice, identify a population equal to or smaller than what can be identified with the PRS at the same CAD risk threshold. Moreover, there is not correlation (max{rho} : 0.137) between PRS and each lipoprotein, indicating that PRS captures different component of CAD etiology and identifies different people at high risk than those identified by lipoproteins, demonstrating to be an invaluable tool in CAD prevention.\n\nOne of the major impairment of the PRS usage in clinical practice is the computational complexity needed to calculate per-individual PRSs. Deep bioinformatics expertise is required to run the entire pipeline, from imputing genomic data, through quality control to result visualisation. For these reasons we developed a Software as a Service (SaaS) for genomic risk prediction of complex diseases. The SaaS is fully automated, GDPR complaint and has been certified as a CE marked medical device. We made the SaaS freely available for research purposes. Researchers willing to use the SaaS can contact research@genomicriskscore.io

7
Effects of pathogenic CNVs on biochemical markers: a study on the UK Biobank

Bracher-Smith, M.; Kendall, K. M.; Rees, E.; Einon, M.; O'Donovan, M.; Owen, M. J.; Kirov, G.

2019-08-06 genomics 10.1101/723270 medRxiv
Top 0.1%
19.2%
Show abstract

BackgroundPathogenic copy number variants (CNVs) increase risk for medical disorders, even among carriers free from neurodevelopmental disorders. The UK Biobank recruited half a million adults who provided samples for biochemical and haematology tests which have recently been released. We wanted to assess how the presence of pathogenic CNVs affects these biochemical test results. MethodsWe called all CNVs from the Affymetrix microarrays and selected a set of 54 CNVs implicated as pathogenic (including their reciprocal deletions/duplications) and present in five or more persons. We used linear regression analysis to establish their association with 28 biochemical and 23 haematology tests. ResultsWe analysed 421k participants who passed our CNV quality control filters and self-reported as white British or Irish descent. There were 268 associations between CNVs and biomarkers that were significant at a false discovery rate of 0.05. Deletions at 16p11.2 had the highest number of significant associations, but several rare CNVs had higher effect sizes indicating that the lack of significance was likely due to the reduced statistical power for rarer events. The distribution of values can be visualised on our interactive website: http://kirov.psycm.cf.ac.uk/. ConclusionsCarriers of many pathogenic CNVs have changes in biochemical and haematology tests, and many of those are associated with adverse health consequences. These changes did not always correlate with increases in diagnosed medical disorders in this population. Carriers should have regular blood tests in order to identify and treat adverse medical consequences early. Levels of cholesterol and related lipids were unexpectedly lower in carriers of CNVs associated with increased weight gain, most likely due to the use of statins by such people.

8
Analyzing Performance of Twist Bioscience Exome Enrichment with Spike-in CNV Backbone Panels at Various Probe Densities Leveraging Golden Helix VS-CNV Analysis Software

Fortier, N.; Rudy, G.; Han, T.; Davassi, A.; Scherer, A.

2024-05-21 bioinformatics 10.1101/2024.05.19.594885 medRxiv
Top 0.1%
18.8%
Show abstract

Clinical Whole Exome Sequencing (WES) offers a high diagnostic yield test by detecting pathogenic variants in all coding genes of the human genome. WES is poised to consolidate multiple genetic tests by accurately identifying Copy Number Variation (CNV) events, typically necessitating microarray analysis. However, standard commercial exome kits are typically limited to targeting exon coding regions, leaving significant gaps in coverage between genes, which could hinder comprehensive CNV detection. To convert microarray CNV calling with NGS, advances in both assay design and computational methods are needed. Addressing the need for comprehensive coverage, Twist Bioscience has developed an enhanced Exome 2.0 Plus Comprehensive Exome Spike-in panel with added CNV "backbone" probes. These probes target common SNPs polymorphic in multiple populations and are evenly distributed in the intergenic and intronic regions, with three varying densities at 25 kb, 50 kb, and 100 kb intervals from highest to lowest resolution respectively. Concurrently, Golden Helix has developed a multi-modal CNV caller designed specifically for target-capture NGS data to detect single-exon to whole-chromosome aneuploidy CNV events. This study evaluates the combined efficacy of the backbone-probe enhanced exome capture kit and VS-CNV 2.6 in identifying known CNVs using the Coriell CNVPANEL01 reference set. The integration of the enhanced capture kit with VS-CNV 2.6 achieved a 100% sensitivity rate for the detection of known CNV events at all three probe densities. The application of best-practice quality metrics, annotations, and filters was shown to have a minimal impact on this high sensitivity. These findings underscore the potential of the augmented Twist Exome in tandem with the VS-CNV caller and VarSeqs annotation and filtering capabilities. This combination presents a promising alternative to conventional microarray assays, potentially consolidating WES and CNV into a single assay obviating the need for additional testing in clinical CNV detection. The studys results advocate for the implementation of this integrated approach as a more efficient and equally sensitive method for CNV analysis in a clinical setting.

9
A novel computational approach to identify cancer cells in scRNA-seq data

Gasper, W. K.; Rossi, F.; Ligorio, M.; Ghersi, D.

2022-04-30 bioinformatics 10.1101/2022.04.28.489880 medRxiv
Top 0.1%
18.5%
Show abstract

Single-cell RNA-seq is an invaluable research tool that allows for the investigation of gene expression in heterogeneous cancer cell populations in ways that bulk RNA-seq cannot. However, normal (i.e., non tumor) cells in cancer samples have the potential to confound the downstream analysis of single-cell RNA-seq data. Several existing methods for identifying tumor cells use copy number variation inference. This work aims to extend existing approaches for identifying cancer cells in single-cell RNA-seq samples by incorporating putative driver alterations. We found that putative driver alterations can be detected in single-cell RNA-seq data and that a subset of cells in tumor samples are enriched in putative driver alterations as compared to normal cells. Furthermore, we show that the number of putative driver alterations and inferred copy number variation are not correlated in all samples. Taken together, our findings suggest that combining copy number variation inference with putative driver mutation load can augment the number of tumor cells that can be confidently included in downstream analyses of single-cell RNA-seq datasets.

10
Automated prediction of the clinical impact of structural copy number variations

Gaziova, M.; Pos, O.; Krampl, W.; Kubiritova, Z.; Kucharik, M.; Radvanszky, J.; Budis, J.; Szemes, T.

2020-07-31 bioinformatics 10.1101/2020.07.30.228601 medRxiv
Top 0.1%
18.3%
Show abstract

Introduction: Copy number variants (CNVs) play an important role in many biological processes, including the development of genetic diseases, making them attractive targets for genetic analyses. The interpretation of the effect of structural variants is a challenging problem due to highly variable numbers of gene, regulatory or other genomic elements affected by the CNV. This led to the demand for the interpretation tools that would relieve researchers, laboratory diagnosticians, genetic counselors, and clinical geneticists from the laborious process of annotation and classification of CNVs. Materials and Methods: We designed a classifier method based on the annotations of CNVs from several publicly available databases. The attributes take into account gene elements, regulatory elements affected by the CNV, as well as other CNVs with known clinical significance that overlap the candidate CNV. We also describe the process of model selection and the construction of training, validation, and test set. Results: The presented approach achieved more than 98% prediction accuracy on both copy number loss and copy number gain variants and can be improved by imposing probability thresholds to eliminate low confidence predictions. Discussion: Method has shown considerable performance in predicting the clinical impact of CNVs and therefore has a great potential to guide users to more precise conclusions. The CNV annotation and pathogenicity prediction can be fully automated, relieving users of tedious interpretation processes. Availability and Implementation: The results can be reproduced by following instructions at {{https://github.com/tsladecek/isv}}.

11
COVID-19: Variant screening, an important step towards precision epidemiology

Chattopadhyay, A.; Lu, T.-P.; Shih, C.-Y.; Lai, L.-C.; Tsai, M.-H.; Chuang, E. Y.

2020-10-19 genomics 10.1101/2020.10.19.345140 medRxiv
Top 0.1%
15.6%
Show abstract

Precision epidemiology using genomic technologies allows for a more targeted approach to COVID-19 control and treatment at individual and population level, and is the urgent need of the day. It enables identification of patients who may be at higher risk than others to COVID-19-related mortality, due to their genetic architecture, or who might respond better to a COVID-19 treatment. The COVID-19 virus, similar to SARS-CoV, uses the ACE2 receptor for cell entry and employs the cellular serine protease TMPRSS2 for viral S protein priming. This study aspires to present a multi-omics view of how variations in the ACE2 and TMPRSS2 genes affect COVID-19 infection and disease progression in affected individuals. It reports, for both genes, several variant and gene expression analysis findings, through (i) comparison analysis over single nucleotide polymorphisms (SNPs), that may account for the difference of COVID-19 manifestations among global sub-populations; (ii) calculating prevalence of structural variations (copy number variations (CNVs) / insertions), amongst populations; and (iii) studying expression patterns stratified by gender and age, over all human tissues. This work is a good first step to be followed by additional studies and functional assays towards informed treatment decisions and improved control of the infection rate.

12
Machine learning approach to assess the pathogenicity of BRCA1/2 genetic variants : brca-NOVUS

Vatsyayan, A.; Scaria, V.

2023-10-20 health informatics 10.1101/2023.10.20.23297295 medRxiv
Top 0.1%
15.4%
Show abstract

Breast cancer is globally the leading type of cancer in terms of both incidence and mortality. BRCA1 and BRCA2 gene variants have long been linked to and studied in context of the disease. Rapid variant discovery has further been made freely accessible by advances in Next-generation sequencing, making it a demanding task to accurately interpret these variants for clinical and research applications. To establish the nature of these variants, the American College of Medical Genetics and Genomics and the Association of Molecular Pathologists (ACMG-AMP) have issued a set of guidelines for variant classification. However, given the huge number of variants associated with the two large and well-studied genes, functional studies or ACMG-AMP classification is a mountainous challenge. Here we describe brca-NOVUS, a machine learning approach trained on a gold-standard ACMG-qualified dataset for the accurate interpretation of variants at large scale. Using two independent test and validation datasets of ACMG-qualified variants, we show that brca-NOVUS can be used to for the classification of variants in clinical as well as research settings.

13
ODDB: Ocular Disease Database for integrated analysis of ocular disease gene drug relationships

Seemab, U.; Kalatanova, A.; Hassan, S.; Tanoli, Z.; Leinonen, H. O.

2025-10-29 bioinformatics 10.1101/2025.10.28.685137 medRxiv
Top 0.1%
15.0%
Show abstract

Ocular diseases such as age-related macular degeneration, glaucoma, diabetic retinopathy, and inherited retinal dystrophies are leading causes of vision loss worldwide, yet existing databases often address only limited aspects of these disorders. To fill this gap, we developed the Ocular Disease Database (ODDB), a web-based resource that integrates genes, biomarkers, variants, and drugs associated to ocular diseases. Data were systematically collected through literature mining of PubMed-indexed journals, the NCBI Gene Expression Omnibus (GEO), and drug regulatory agency datasets. Multi-omics, experimental, and clinical information were harmonized using standardized integration workflows. The database is organized according to two complementary ontologies: one based on the anatomical site of pathology (cornea, retina, optic nerve) and another on gene inheritance pattern. ODDB currently covers over 170 ocular diseases, more than 1190 genes, 2400+ variants, and 386 drugs, including both approved and investigational compounds. Each record includes detailed annotations of associated genes, variants, therapeutic targets, and mechanisms of action. The platform supports interactive querying and network-based visualization of disease-gene-drug relationships. All data was internally validated for accuracy and are compliant with FAIR principles, ensuring accessibility and interoperability. ODDB (https://www.oculardiseases.fi/) provides a comprehensive and standardized reference for exploring molecular mechanisms and therapeutic opportunities in ocular diseases.

14
Determination of Copy Number Variations and Affected Gene Networks in Breast Cancer from Mexican Patients

Larios-Serrato, V.; Valdez-Salazar, H.-A.; Torres, J.; Camorlinga-Ponce, M.; Pina-Sanchez, P.; Mayani-Viveros, H.; Ruiz-Tachiquin, M.-E.

2025-08-31 genomics 10.1101/2025.08.28.672450 medRxiv
Top 0.1%
14.8%
Show abstract

Triple-negative breast cancer (TNBC) is an aggressive subtype with limited treatment options and high molecular heterogeneity. In this study, we performed a genome-wide analysis of copy number variations (CNVs) using high-density microarrays in tumor tissue (TUM), tumor adjacent tissue (ADJ), and leukocytes (LEU) from five Mexican TNBC patients. We identified both unique and shared CNVs across tissues, including alterations in key chromosomal regions such as 1q23.3, 1q32.1, and 8q24.3, which harbor oncogenes like MYC, MCL1, and BCL9. Losses in 6q25.2 affecting ESR1 were also detected. CNVs were enriched in genes related to the Hallmarks of Cancer, with TUM samples showing profiles associated with proliferation, metastasis, and immune evasion; ADJ samples with growth suppression; and LEU samples with genomic instability. Pathway enrichment analyses revealed disrupted functions in DNA repair, extracellular matrix organization, and TP53 signaling in TUM. Notably, EGFR, ERCC4, and HSP90AB1 emerged as central nodes in interaction networks and may serve as markers or therapeutic targets. This is the first CNV profiling study of its kind in TNBC from Mexican patients, highlighting the importance of including underrepresented populations in genomic research to uncover distinct molecular signatures and potential diagnostic or therapeutic avenues. ImplicationsMolecular signatures of breast cancer (BC), predicted bioinformatically, involve common and distinct CNV-Hallmarks of Cancer genes, which are suitable candidates for screening as potential BC markers.

15
Towards Analyzing the Reclassification Dynamics of ClinVar Variants

Ruscheinski, A.; Reimler, A. L.; Ewald, R.; Uhrmacher, A. M.

2023-01-25 bioinformatics 10.1101/2023.01.24.525342 medRxiv
Top 0.1%
13.2%
Show abstract

ClinVar aggregates information about genetic variants and their relation to human diseases and forms a valuable resource for clinical diagnostics. The assessment of variants, e.g., as being benign or pathogenetic, changes over time. We collected variant classification histories of ClinVar releases and used those to derive discrete-time Markov chains to investigate the reclassification dynamics of variants in ClinVar in terms of transition probabilities between ClinVar releases and reclassifications.

16
A broad exome study of the genetic architecture of asthma reveals novel patient subgroups

Cameron-Christie, S.; Mackay, A.; Wang, Q.; Olsson, H.; Angermann, B.; Lassi, G.; Lindgren, J.; Hühn, M.; Cameron-Christie, Y. O.; Gavala, M.; Wang, J.; Povysil, G.; Deevi, S. V. V.; Belfield, G.; Dillmann, I.; Muthas, D.; Cohen, S.; Young, S.; Platt, A.; Petrovski, S.

2020-12-11 genomics 10.1101/2020.12.10.419663 medRxiv
Top 0.1%
13.1%
Show abstract

IntroductionAsthma risk is a complex interplay between genetic susceptibility and environment. Despite many significantly-associated common variants, the contribution of rarer variants with potentially greater effect sizes has not been as extensively studied. We present an exome-based study adopting 24,576 cases and 120,530 controls to assess the contribution of rare protein-coding variants to the risk of early-onset or all-comer asthma. MethodsWe performed case-control analyses on three genetic units: variant-, gene- and pathway-level, using sequence data from the Scandinavian Asthma Genetic Study and UK Biobank participants with asthma. Cases were defined as all-comer asthma (n=24,576) and early-onset asthma (n=5,962). Controls were 120,530 UK Biobank participants without reported history of respiratory illness. ResultsVariant-level analyses identified statistically significant variants at moderate-to-common allele frequency, including protein-truncating variants in FLG and IL33. Asthma risk was significantly increased not only by individual, common FLG protein-truncating variants, but also among the collection of rare-to-private FLG protein-truncating variants (p=6.8x10-7). This signal was driven by early-onset asthma and did not correlate with circulating eosinophil levels. In contrast, a single splice variant in IL33 was significantly protective (p=8.0x10-10), while the collection of remaining IL33 protein-truncating variants showed no class effect (p=0.54). A pathway-based analysis identified that protein-truncating variants in loss-of-function intolerant genes were significantly enriched among individuals with asthma. ConclusionsAccess to the full allele frequency spectrum of protein-coding variants provides additional clarity about the potential mechanisms of action for FLG and IL33. Beyond these two significant drivers, we detected a significant enrichment of protein-truncating variants in loss-of-function intolerant genes.

17
Proteomic Fingerprinting: A novel privacy concern

Hill, A. C.; Litkowski, E. M.; Manichaikul, A.; Lange, L.; Pratte, K. A.; Kechris, K. J.; DeCamp, M.; Coors, M.; Ortega, V. E.; Rich, S. S.; Rotter, J. I.; Gerzsten, R. E.; Clish, C. B.; Curtis, J.; Hu, X.; Ngo, D.; ONeal, W. K.; Meyers, D.; Bleecker, E.; Hobbs, B. D.; Cho, M. H.; Banaei-kashani, F.; Bowler, R. P.

2022-04-12 genetic and genomic medicine 10.1101/2022.04.06.22269907 medRxiv
Top 0.1%
13.0%
Show abstract

IntroductionPrivacy protection is a core principle of genomic research but needs further refinement for high-throughput proteomic platforms. MethodsWe identified independent single nucleotide polymorphism (SNP) quantitative trait loci (pQTL) from COPDGene and Jackson Heart Study (JHS) and then calculated genotype probabilities by protein level for each protein-genotype combination (training). Using the most significant 100 proteins, we applied a naive Bayesian approach to match proteomes to genomes for 2,812 independent subjects from COPDGene, JHS, SubPopulations and InteRmediate Outcome Measures In COPD Study (SPIROMICS) and Multi-Ethnic Study of Atherosclerosis (MESA) with SomaScan 1.3K proteomes and also 2,646 COPDGene subjects with SomaScan 5K proteomes (testing). We tested whether subtracting mean genotype effect for each pQTL SNP would obscure genetic identity. ResultsIn the four testing cohorts, we were able to correctly match 90%-95% their proteomes to their correct genome and for 95%-99% we could match the proteome to the 1% most likely genome. With larger profiling (SomaScan 5K), correct identification was > 99%. The accuracy of matching in subjects with African ancestry was lower ([~]60%) unless training included diverse subjects. Mean genotype effect adjustment reduced identification accuracy nearly to random guess. ConclusionLarge proteomic datasets (> 1,000 proteins) can be accurately linked to a specific genome through pQTL knowledge and should not be considered deidentified. These findings suggest that large scale proteomic data be given privacy protections of genomic data, or that bioinformatic transformations (such as adjustment for genotype effect) should be applied to obfuscate identity.

18
Tracing evolution history of 100 whole genome sequences of diffuse stem brain tumor

Zou, L.

2021-03-16 genomics 10.1101/2021.03.14.435347 medRxiv
Top 0.1%
12.9%
Show abstract

Diffuse intrinsic pontine glioma (DIPG) is a deadly disease among young children. The evolution path and mutational processes giving rise to DIPG remain elusive. We analyzed 100 whole genome sequences (WGS) from 60 DIPG patients. This revealed 25% DIPGs acquired whole-genome duplications (WGD) early during tumor evolution. WGD samples are associated with loss of TP53 and poorer survival. In addition, almost all WGD samplers harbor complex structural variations (SVs) and show characteristic short microhomology at SV breakpoints. Mutation analysis revealed that H3K27M driver mutation is acquired early during tumor clonal evolution. Mutation signature analysis identified a unique mutational process at a late stage of tumor evolution. This study revealed that tumor evolution of DIPG is characterized by chromosomal instability shaped by DNA repair defects and dynamic mutational processes. Our work shed new insights on the disease pathogenesis of DIPG and provided rationale for designing novel therapy for this deadly disease.

19
ClinCNV: multi-sample germline CNV detection in NGS data

Demidov, G.; Sturm, M.; Ossowski, S.

2022-06-13 bioinformatics 10.1101/2022.06.10.495642 medRxiv
Top 0.1%
12.8%
Show abstract

Germline copy number variants (CNVs) are a common source of genomic variation involved in many genetic disorders, and their detection is crucial for clinical molecular diagnostics. Genomic microarrays, quantitative polymerase chain reaction (qPCR), and multiplex ligation-dependent probe amplification (MLPA) have been widely used for CNV detection in clinics for many years. Similarly, next-generation sequencing (NGS) applications such as whole-genome sequencing (WGS) and whole-exome sequencing (WES) are well-established, highly accurate techniques for the detection of single nucleotide variants (SNVs) and small insertions and deletions (indels). However, CNV detection using NGS remains challenging due to short read lengths, smaller than CNVs sizes. CNV detection using read coverage depths summarized in genomic regions is affected by various biases that arise during the library preparation and sequencing. We have developed a novel strategy for detecting CNVs, implemented in the tool ClinCNV (freely available on https://github.com/imgag/ClinCNV). ClinCNV does multi-sample normalization and CNV calling, using an original algorithm taking the best from the circular binary segmentation method and Hidden Markov model-based approaches. Here, we describe the methods and discuss the results obtained by applying ClinCNV to thousands of clinical WES, WGS, and shallow-WGS samples in various clinical and research settings.

20
Spatial Distribution of Missense Variants within Complement Proteins Associates with Age Related Macular Degeneration

Grunin, M.; de Jong, S.; Palmer, E. L.; Jin, B.; Rinker, D.; Moth, C.; Capra, J. A.; Haines, J. L.; Bush, W.; den Hollander, A.; International Age-related Macular Degeneration Genomics Consortium,

2023-08-31 genetic and genomic medicine 10.1101/2023.08.28.23294686 medRxiv
Top 0.1%
12.8%
Show abstract

PurposeGenetic variants in complement genes are associated with age-related macular degeneration (AMD). However, many rare variants have been identified in these genes, but have an unknown significance, and their impact on protein function and structure is still unknown. We set out to address this issue by evaluating the spatial placement and impact on protein structureof these variants by developing an analytical pipeline and applying it to the International AMD Genomics Consortium (IAMDGC) dataset (16,144 AMD cases, 17,832 controls). MethodsThe IAMDGC dataset was imputed using the Haplotype Reference Consortium (HRC), leading to an improvement of over 30% more imputed variants, over the original 1000 Genomes imputation. Variants were extracted for the CFH, CFI, CFB, C9, and C3 genes, and filtered for missense variants in solved protein structures. We evaluated these variants as to their placement in the three-dimensional structure of the protein (i.e. spatial proximity in the protein), as well as AMD association. We applied several pipelines to a) calculate spatial proximity to known AMD variants versus gnomAD variants, b) assess a variants likelihood of causing protein destabilization via calculation of predicted free energy change (ddG) using Rosetta, and c) whole gene-based testing to test for statistical associations. Gene-based testing using seqMeta was performed using a) all variants b) variants near known AMD variants or c) with a ddG >|2|. Further, we applied a structural kernel adaptation of SKAT testing (POKEMON) to confirm the association of spatial distributions of missense variants to AMD. Finally, we used logistic regression on known AMD variants in CFI to identify variants leading to >50% reduction in protein expression from known AMD patient carriers of CFI variants compared to wild type (as determined by in vitro experiments) to determine the pipelines robustness in identifying AMD-relevant variants. These results were compared to functional impact scores, ie CADD values > 10, which indicate if a variant may have a large functional impact genomewide, to determine if our metrics have better discriminative power than existing variant assessment methods. Once our pipeline had been validated, we then performed a priori selection of variants using this pipeline methodology, and tested AMD patient cell lines that carried those selected variants from the EUGENDA cohort (n=34). We investigated complement pathway protein expression in vitro, looking at multiple components of the complement factor pathway in patient carriers of bioinformatically identified variants. ResultsMultiple variants were found with a ddG>|2| in each complement gene investigated. Gene-based tests using known and novel missense variants identified significant associations of the C3, C9, CFB, and CFH genes with AMD risk after controlling for age and sex (P=3.22x10-5;7.58x10-6;2.1x10-3;1.2x10-31). ddG filtering and SKAT-O tests indicate that missense variants that are predicted to destabilize the protein, in both CFI and CFH, are associated with AMD (P=CFH:0.05, CFI:0.01, threshold of 0.05 significance). Our structural kernel approach identified spatial associations for AMD risk within the protein structures for C3, C9, CFB, CFH, and CFI at a nominal p-value of 0.05. Both ddG and CADD scores were predictive of reduced CFI protein expression, with ROC curve analyses indicating ddG is a better predictor (AUCs of 0.76 and 0.69, respectively). A priori in vitro analysis of variants in all complement factor genes indicated that several variants identified via bioinformatics programs PathProx/POKEMON in our pipeline via in vitro experiments caused significant change in complement protein expression (P=0.04) in actual patient carriers of those variants, via ELISA testing of proteins in the complement factor pathway, and were previously unknown to contribute to AMD pathogenesis. ConclusionWe demonstrate for the first time that missense variants in complement genes cluster together spatially and are associated with AMD case/control status. Using this method, we can identify CFI and CFH variants of previously unknown significance that are predicted to destabilize the proteins. These variants, both in and outside spatial clusters, can predict in-vitro tested CFI protein expression changes, and we hypothesize the same is true for CFH. A priori identification of variants that impact gene expression allow for classification for previously classified as VUS. Further investigation is needed to validate the models for additional variants and to be applied to all AMD-associated genes.